Papers with aggregated metrics
Semantic Diversity for Natural Language Understanding Evaluation in Dialog Systems (2020.coling-industry)
Copied to clipboard
| Challenge: | a dialog system is used to evaluate NLU models using aggregated metrics on a large number of utterances. |
| Approach: | They propose a method to generate a test set with high semantic diversity for NLU evaluation in dialog systems. |
| Outcome: | The proposed test sets are based on high diversity of utterances from different regions of the utteration embedding space. |
REMIND: Memorization and Unlearning in LLMs Through the Lens of Input Loss Landscapes (2026.acl-long)
Copied to clipboard
| Challenge: | REMIND is a framework that diagnoses residual memorization states by probing local ILL curvature over semantically coherent neighborhoods. |
| Approach: | They propose a framework that diagnoses memorization states by probing local ILL curvature over semantically coherent neighborhoods. |
| Outcome: | The proposed framework outperforms baseline models with 82% multi-class ROC-AUC and 2 higher AUC at 1% FPR. |